Papers with automatic classification

11 papers
Detecting Primary Progressive Aphasia (PPA) from Text: A Benchmarking Study (2026.findings-eacl)

Copied to clipboard

Challenge: Primary progressive aphasia (PPA) is a neurodegenerative disorder characterized by progressive language deficits as the primary symptom.
Approach: They benchmarked the performance of traditional machine learning models with various feature extraction techniques, transformer-based models, and large language models (LLMs) they found that transformer-Based models exceeded chance-level performance in terms of balanced accuracy, while MLP using MentalBert’s embeddings achieved the highest accuracy.
Outcome: The proposed models outperform chance-level models in terms of balanced accuracy while using MentalBert’s embeddings achieve the highest accuracy.
Annotation and Classification of Relevant Clauses in Terms-and-Conditions Contracts (2024.lrec-main)

Copied to clipboard

Challenge: Using Large Language Models (LLMs) as foundational models, we propose a new annotation scheme to classify different types of clauses in Terms-and-Conditions contracts.
Approach: They propose to use a new annotation scheme to classify clauses in Terms-and-Conditions contracts to support legal experts in identifying and assessing problematic issues.
Outcome: The proposed annotation scheme achieves accuracies ranging from .79 to .95 on validation tasks.
Multimodal Pipeline for Collection of Misinformation Data from Telegram (2022.lrec-1)

Copied to clipboard

Challenge: a large portion of misinformation is spread via multimodal means, such as images and videos . a new pipeline for collecting misinformation from Telegram allows us to collect a greater variety of mis-information examples .
Approach: They propose to use AI to understand misinformation flow across social media platforms . they collect data from Telegram groups which promote COVID-19 misinformation .
Outcome: The proposed dataset contains almost one million messages from 2k different public channels related to spreading COVID-19 misinformation.
The Automatic Annotation of the Semiotic Type of Hand Gestures in Obama’ s Humorous Speeches (L18-1)

Copied to clipboard

Challenge: Existing studies on hand gestures from video-recorded speeches have not identified them.
Approach: They annotated and analysed hand gestures produced by Barack Obama . they trained machine learning algorithms to classify the semiotic type of hand gesture .
Outcome: The proposed method can be used to classify hand gestures on video-recorded speeches and in advanced multimodal interactive systems.
Identifying the Human Values behind Arguments (2022.acl-long)

Copied to clipboard

Challenge: et al., 2003) examines human values in natural language arguments . authors provide a dataset of 5270 arguments from four geographical cultures .
Approach: They propose a multi-level taxonomy of human values with 54 values and a dataset of 5270 arguments from four geographical cultures, manually annotated for human values.
Outcome: The proposed model shows that human values are more diverse than previously thought . it shows that people disagree on the best course forward on controversial issues .
Classifying Referential and Non-referential It Using Gaze (D18-1)

Copied to clipboard

Challenge: a particular problem for anaphora resolution systems is the pronoun it, which can be used both referentially and non-referentially.
Approach: They use eye-tracking data to learn how humans perform disambiguation and use it to improve automatic classification.
Outcome: The proposed system outperforms a baseline and outperformed linguistic-based approaches.
Classification of Closely Related Sub-dialects of Arabic Using Support-Vector Machines (L18-1)

Copied to clipboard

Challenge: Existing studies on dialect identification have focused on binary classifications between colloquial Arabic and dialectal Egyptian .
Approach: They propose to use an n-gram based SVM to classify on a fine-grained sub-dialectal level and compare it to methods used in dialect classification such as vocabulary pruning.
Outcome: The proposed method is compared to methods used in dialect classification such as vocabulary pruning of shared items across dialects.
A Corpus for Suggestion Mining of German Peer Feedback (2022.lrec-1)

Copied to clipboard

Challenge: e.g. Massive Open Online Courses (MOOCs) are increasingly important to meet the demand for feedback in large scale classes.
Approach: They propose to use peer feedback to detect suggestions on how to improve the work of students in a german university course.
Outcome: The proposed corpus is the first student peer feedback corpus in germany and has been labelled with a new annotation scheme.
Hierarchical Multi-Label Classification of Scientific Documents (2022.emnlp-main)

Copied to clipboard

Challenge: Automated topic classification is a useful tool for managing scientific documents in a digital collection.
Approach: They propose a hierarchical multi-label text classification dataset with keyword labeling as an auxiliary task.
Outcome: The proposed model achieves a Macro-F1 score of 34.57% and is publicly available.
Invisible to People but not to Machines: Evaluation of Style-aware HeadlineGeneration in Absence of Reliable Human Judgment (2020.lrec-1)

Copied to clipboard

Challenge: Using a data alignment strategy and different training/testing settings, we aim at decoupling content from style and preserving the latter in generation.
Approach: They propose a fine-grained evaluation strategy based on automatic classification to evaluate generated headlines' quality in terms of their newspaper-compliance.
Outcome: The proposed model learns newspaper-specific style, but humans aren't reliable judges for this task, and deserves particular care in its design.
Murre24: Dialect Identification of Finnish Internet Forum Messages (2024.lrec-main)

Copied to clipboard

Challenge: 94 million messages posted on the largest Finnish internet forum, Suomi24, are classified to present either the standard language, one of the seven traditional dialects, a colloquial style or the Helsinki slang.
Approach: They present a collection of dialectal messages posted on the largest Finnish internet forum, Suomi24 . they manually annotated a dataset and used it to train dialect identification models .
Outcome: The proposed method is the best for differentiating standard Finnish from non-standard Finnish, while fine-tuning a BERT-based model achieves best scores on the final dialect identification task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations